Back

BMC Medical Research Methodology

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match BMC Medical Research Methodology's content profile, based on 47 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.

1
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 0.1%
19.6%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.

2
NeMMo: an improved statistical algorithm for excess all-cause mortality surveillance and monitoring

Lytras, T.; Athanasiadou, M.

2026-08-17 epidemiology 10.64898/2026.08.14.26360477 medRxiv
Top 0.1%
18.8%
Show abstract

Background: Reliable estimation of excess mortality is central to population health surveillance. We introduce NeMMo (New Mortality Model), an evolution of the EuroMOMO model for estimating weekly all-cause expected mortality, and assess its behaviour and performance on empirical data. Methods: NeMMo incorporates population offsets, stratifies observed deaths by age group and models seasonality using a periodic B-spline rather than a Serfling-type sinusoidal function. Baseline weeks are selected by a data-driven procedure minimizing the skewness of the residuals before refitting the model, instead of relying solely on fixed calendar windows. NeMMo enables pooling across age groups, direct age standardization and incorporation of external predictors. We applied NeMMo and EuroMOMO to mortality and population data downloaded from Eurostat for 31 countries from 2015 onwards, excluding the COVID-19 pandemic period from baseline estimation. Results: For most countries NeMMo produced a higher expected mortality baseline that better tracked observed deaths, as well as tighter prediction intervals and higher maximum Z-scores, suggesting improved discrimination of mortality excesses. Z-scores and P-scores during non-pandemic weeks were closer to zero with NeMMo than with EuroMOMO but further elevated during pandemic weeks, providing greater separation between pandemic and non-pandemic mortality. Incorporating population offsets resulted in negative linear trends across all countries, consistent with declining mortality after accounting for demographic changes. The periodic B-spline identified substantial heterogeneity in the shape and timing of seasonal mortality that was not captured by a sinusoidal function. Conclusions: NeMMo provides a flexible and parsimonious framework for all-cause mortality surveillance that improves the established EuroMOMO model and offers theoretical, empirical and practical advantages. It is thus suitable both for detecting short-term spikes and for the long-term, age-adjusted quantification and comparison of mortality excesses that has become increasingly important since the COVID-19 pandemic. The accompanying 'nemmo' package for R facilitates its widespread adoption and application.

3
Multi-model LLM assessment of Quality Control Circlemethodological quality: a designed-anchor reliabilitystudy

LIn, H.; Lyu, J.

2026-08-13 health informatics 10.64898/2026.08.12.26360276 medRxiv
Top 0.1%
15.3%
Show abstract

BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.

4
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.1%
15.2%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

5
Algorithmic Ascertainment of Cause of Death from Longitudinal Real-World Medical Claims Data: Development and Validation

McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360606 medRxiv
Top 0.1%
12.7%
Show abstract

This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.

6
Artificial Scientific Intelligence for Measurement-burden-aware Modelling and Interpretation of Multi-site Bone Mineral Density

Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.

2026-09-01 health informatics 10.64898/2026.08.30.26361665 medRxiv
Top 0.1%
12.5%
Show abstract

Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.

7
Towards understanding the disease landscape of clinical trials in Germany: Ontology and embedding-based pipelines versus Large Language Models for ICD-10 Harmonization

Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.

2026-08-06 health informatics 10.64898/2026.08.04.26359616 medRxiv
Top 0.1%
10.7%
Show abstract

Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.

8
Does Data Preprocessing Affect Tree-Based Super Learners? An Investigation of Ensemble Optimization and Oracle Properties in Clinical Classification.

Darko, R.; Dwumah, D.; Agyapong, K. S.; Agyenim-Boateng, Y.; Darko Anim, R.; Wisdom Jakper, J.; Owusu-Ansah, N. K.; Owusu-Ansah, R.

2026-08-24 health informatics 10.64898/2026.08.20.26360880 medRxiv
Top 0.1%
9.7%
Show abstract

Machine learning workflows frequently incorporate data preprocessing to enhance predictive performance. However, the need for Super Learner ensembles made up only of preprocessing-invariant tree-based algorithms remains unexplored. Using three benchmark clinical classification datasets, this study examined how preprocessing affected the Super Learner's prediction performance, learner weight distribution, and oracle behavior. The Heart Disease (207 observations), Indian Liver Patient Dataset (583 observations), and Pima Indians Diabetes (768 observations) datasets were used to create a Super Learner ensemble model that included Classification and Regression Trees (CART), Random Forest, Ranger, and Extreme Gradient Boosting (XGBoost). Models were evaluated under raw and preprocessed data conditions using repeated cross-validation. Predictive performance was assessed using the area under the receiver operating characteristic curve (AUC), Matthews correlation coefficient (MCC), and Brier score. Learner weight allocation and Oracle Gap were compared using paired Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment. Preprocessing produced negligible changes in predictive performance for the Heart Disease and Pima datasets. For the ILPD dataset, preprocessing significantly improved AUC (0.746 to 0.752; adjusted p = 0.0017) and reduced the Brier score (0.177 to 0.175; adjusted p < 0.001). Learner weights remained largely stable, although Random Forest replaced Ranger as the dominant learner for the Heart Disease dataset. Oracle Gaps remained extremely small (<0.002) across all datasets and did not differ significantly between preprocessing conditions. Preprocessing provides limited benefit for Super Learner ensembles composed of preprocessing-invariant learners and does not materially alter their oracle behavior. Preprocessing decisions should therefore be guided by dataset characteristics rather than adopted as a universal modelling practice.

9
From surveillance maturity to analytical readiness: an estimand-first framework for real-time outbreak analysis under imperfect data

Verheyden, J. G. L.; Mudogo, C. N.; Jacquet, W.

2026-08-14 epidemiology 10.64898/2026.08.12.26360299 medRxiv
Top 0.1%
7.9%
Show abstract

Background: Real-time outbreak analyses are often requested before surveillance systems have stabilised or epidemics have generated enough information for the desired inference. Existing approaches address surveillance quality, forecasting, estimands and identifiability separately, but do not provide a common rule for deciding which analytical product is supportable at a particular data vintage. We developed an estimand-first framework for analytical readiness. Methods: The framework distinguishes surveillance maturity (S), epidemic-process informativeness (E) and estimand-specific analytical readiness, defined as whether the available data vintage, observation process, method and decision-matched validation support a specified inference for a specified decision. We stress-tested four implications using longitudinal data from the 2018-2020 Ebola response in eastern Democratic Republic of the Congo (DRC), archived geographic forecasts, independent forecasting data from Western Area, Sierra Leone, and a targeted mortality-identifiability experiment. Results: During a documented DRC surveillance disruption and recovery, seven-day persistence forecasts had all-health-zone absolute errors of 1, 3, 18 and 4 cases across pre-shock, acute-shock, early-recovery and recovery origins; the largest error occurred during early recovery. Four-week reported-case trend multipliers changed from 0.67 and 0.73 to 1.12 and 1.29, while the final fit was strongly overdispersed (Pearson dispersion 7.65), demonstrating asynchronous readiness across estimands. Archived geographic forecasts improved a Top-3 allocation decision over cumulative burden at only one origin despite consistently lower Brier scores for one specification. In Western Area, persistence forecast mean absolute error increased from 39.1 cases at one week to 82.3 at four weeks, and a history-to-horizon ratio did not define a universal threshold. An observed reported case-fatality ratio of 0.40 was compatible with constructed latent fatality values from 0.10 to 0.80; increasing the reported denominator narrowed sampling uncertainty without reducing structural uncertainty. Conclusions: Analytical readiness is task- and vintage-specific rather than a property of a dataset. More data, model convergence or narrow intervals cannot substitute for estimand definition, observation-process awareness, decision-matched validation and explicit identification analysis. Keywords: outbreak analytics; surveillance maturity; analytical readiness; estimand; identifiability; forecasting; Ebola; reporting process; decision-matched validation

10
A software package for simple and rigorous survival machine learning analysis in biomedical research

Pybus, A.; Qiu, J.; Morais Lyra, P. C.; Dang, K.; Narvaez-Bandera, I.; Jolaogun, T.; Goecks, J.

2026-08-10 health informatics 10.64898/2026.08.05.26359034 medRxiv
Top 0.1%
6.8%
Show abstract

Survival analysis is a fundamental technique in biomedical research for modeling time-to-event data. It enables the identification of prognostic factors in disease, compares survival outcomes across treatment groups, and performs targeted treatment selection. A variety of machine learning (ML) approaches to survival analysis have emerged to complement classical statistical methods, especially for high-dimensional datasets with complex, nonlinear interactions between features. However, using survival ML methods requires addressing challenges such as censoring-unaware evaluation, overfitting, selecting performance metrics, and data leakage. To address these and other difficulties in using survival ML models, we developed the mlsurv software package. mlsurv is an open-source Python package built around three major design principles: 1) methodological rigor, including evidence-based model selection, leakage-free pipelines, and multi-metric evaluation, 2) multi-scale evaluation and interpretation, including population and subpopulation evaluation, patient-level explanations, and feature analysis, and 3) automated trust and transparency, including limitation flagging and TRIPOD+AI-aligned reporting. mlsurv bundles ten models spanning linear, ensemble, kernel, and deep learning families within a unified software package. We demonstrate mlsurv on the Chowell immunotherapy cohort (n=1,479). The survival-trained models achieve a test concordance index of 0.73 for overall survival prediction. Further, risk scores strongly correlate with the response-trained LORIS clinical score (|{rho}| up to 0.84), reflecting the overlap between prognostic and predictive signal. mlsurv enables biomedical researchers to conduct rigorous, multi-model survival analysis and benchmarking using minimal code with default best practices rather than implementing custom scripts and methodological safeguards from scratch.

11
Addressing Measurement Error of Machine-Learned Physical Activity in Nonlinear Dose-Response Survival Analysis: Development and Evaluation of Accelerated Failure Time, Spline, and Simulation-Extrapolation Method

Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.

2026-08-31 epidemiology 10.64898/2026.08.25.26361155 medRxiv
Top 0.2%
5.6%
Show abstract

Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.

12
Retrieval-Augmented Large Language Models for Clinically Aligned Adverse Event Coding in Acute Myeloid Leukemia Clinical Trials

Dashti, N.; Schneider, M. M. K.; Eckardt, J. N.; Fiebig, F.; Schweigler, D.; Buttner, S.; Middeke, J. M.; Bornhauser, M.; Rollig, C.; Kather, J. N.; Wiest, I. C.

2026-08-18 health informatics 10.64898/2026.08.17.26360282 medRxiv
Top 0.2%
5.6%
Show abstract

Background: Adverse event (AE) coding is essential for safety monitoring in oncology clinical trials, particularly in acute myeloid leukemia (AML), where intensive therapies are associated with frequent and heterogeneous toxicities requiring standardized MedDRA (Medical Dictionary for Regulatory Activities) coding. However, manual Low-Level Term (LLT) assignment remains labor-intensive, subjective, and difficult to scale. Although large language models (LLMs) have emerged as promising decision-support tools for automated coding, unguided zero-shot generation remains insufficient for reliable fine-grained MedDRA coding. Objective: To develop and evaluate a retrieval-augmented reasoning pipeline for clinically aligned LLT-level MedDRA coding of free-text adverse events from prospective AML clinical trials. Methods: We implemented a retrieval-augmented reasoning pipeline inspired by the retrieval-augmented generation (RAG) paradigm using LLaMA-3.3-70B-Instruct as the primary backbone and benchmarked the framework across multiple open instruction-tuned LLMs. Dense semantic retrieval first generated a constrained top-100 LLT candidate set for each AE, followed by structured LLM reasoning to select a single best-matching LLT and deterministic mapping to Preferred Term (PT) and System Organ Class (SOC) levels. The pipeline was evaluated retrospectively on AE datasets from three prospective AML clinical trials (MOSAIC, DELTA, and DaunoDouble) with automated LLT/PT/SOC metrics and expert-assessed Clinical Correctness Rate (CCR). Results: Clinical expert review showed high clinical acceptability of the RAG pipeline across datasets (91-97%). Under automated evaluation, the pipeline achieved LLT exact accuracy of 50-58%, PT accuracy of 78-85%, and SOC accuracy of 90-93%. Zero-shot generation and random candidate selection performed substantially worse. Semantic retrieval more often included the coder-assigned LLT among the candidate terms available to the model than retrieval based on lexical similarity. Multi-model benchmarking showed that backbone choice mainly affected LLT exact agreement, whereas PT and SOC performance remained comparatively stable. Conclusions: Retrieval-augmented reasoning supports clinically aligned MedDRA coding of free-text adverse events under realistic candidate constraints in AML clinical trials. Evaluation across three AML clinical trials showed that strict LLT-level string agreement underestimated clinical ap-propriateness, highlighting the importance of combining hierarchical evaluation metrics with clini-cal expert validation for AI-assisted MedDRA coding in hematology trials.

13
Describing health inequalities without distortion: Simple-Means MAIHDA vs Random-Effects MAIHDA

Merlo, J.; Bashir, N. Z.; Rodriguez-Lopez, M.; Khalaf, K.; Öberg, J.; Perez-Vicente, R.

2026-08-18 epidemiology 10.64898/2026.08.17.26360592 medRxiv
Top 0.2%
5.2%
Show abstract

Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) describes health inequalities through three components: (i) specific contextual effects (SCE), (ii) general contextual effects (GCE), and (iii) discriminatory accuracy of the context. We present Simple-Means MAIHDA (S-MAIHDA), which estimates each stratum directly from its observed individuals, with no distributional assumption. The observed proportions are unbiased whatever the stratum size, and their confidence intervals report the uncertainty honestly. S-MAIHDA operationalises the three components on the probability scale. The SCE are the raw and standardised stratum prevalences and the modification of the sociodemographic average differences by the area. The GCE are the variance partition coefficient (VPC) and the contextual structuring of the between-stratum inequality, expressed as the contextual clustering of inequalities, the additive sociodemographic differences, and the contextual modification of inequalities (CMI). The contextual discriminatory accuracy is expressed by the area under the ROC curve (AUC), and the sensitivity and specificity at the population prevalence as the threshold for a possible intervention. Because its estimates are the observed data themselves, S-MAIHDA is the canonical description, and the compare diagnostic quantifies how Random-Effects MAIHDA (RE-MAIHDA), the usual implementation, departs from it: RE shrinkage pulls small strata towards the overall mean and can hide the very inequalities the analysis seeks. The approach is implemented in the smaihda Stata command and reproduced in free Python code. We illustrate S-MAIHDA on register data from Malmo, Sweden (43,291 individuals; 300 area-sociodemographic strata), showing how the three components separate two contrasting outcomes: psychotropic medication use, almost purely sociodemographic, stable across areas, with weak contextual structuring (VPC {approx} 4%, CMI {approx} 0%); and choice of a private general practitioner, strongly geographical (VPC {approx} 11%, CMI {approx} 17%), with the sociodemographic differences reshaped and amplified in wealthy areas. RE-MAIHDA attenuated inequalities. For describing inequalities, S-MAIHDA preserves what the data show.

14
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.2%
4.9%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

15
Sharing Aggregated Patient Counts in Place of Line-Level EHR Data: Analytic Fidelity and the Limits of Count Suppression for Privacy

Chen, Y.; McMurry, A.; Gottlieb, D.; Jones, J. R.; Strober, B. J.; Mandl, K. D.

2026-08-21 health informatics 10.64898/2026.08.18.26359984 medRxiv
Top 0.2%
4.6%
Show abstract

Objective. Privacy regulation constrains sharing line-level electronic health records (EHR) across institutions. One alternative is to aggregate counts into a cube, a table of counts for every combination of categorical variables, with cells below a threshold suppressed. This study asked whether common analyses on the cube reproduce conclusions from line-level data, and whether suppression prevents recovery of the small cells it is meant to hide. Materials and Methods. A Bayesian count-inference pipeline was built that reconstructs suppressed counts and doubles as a reconstruction attack. Applied to 285 pediatric kidney-transplant patients at Boston Children's Hospital, statistical fidelity (Jensen-Shannon divergence, Cramer's V, and R2) and analytical utility (marginal distributions, subgroup graft rejection odds ratios, and logistic-regression classification) were evaluated. Conditional Tabular GAN (CTGAN) synthetic data served as a comparator. Results. Statistical analyses on the cube recapitulated results from line-level data. Across 106 demographic-by-medication subgroups, a bootstrap mean of 3.5 subgroups showed a significant graft-rejection association. The cube's odds-ratio sign changes reversed no significant associations, versus 2.3 for CTGAN. The same reconstruction also defeated suppression: in a 10-variable cube, 76.6% of suppressed cube cells were recovered exactly (14,554 of 18,994), including 85.5% of single-patient cells. Discussion. The cube reproduced common kidney-transplant analyses, but the same reconstruction also recovered suppressed cells; fidelity and privacy risk are thus two faces of one reconstruction rather than independent properties. Conclusions. The cube is a useful surrogate for these kidney-transplant analyses only when paired with a stronger privacy mechanism. This study demonstrated reconstructability of suppressed counts, not re-identification.

16
Systematic Data Fitness Assessment Improves Validity and Replicability of Research Using Real-World Data

Razzaghi, H.; Wieand, K.; Pinkney, A.; Bailey, C.

2026-08-10 epidemiology 10.64898/2026.08.05.26359818 medRxiv
Top 0.2%
4.5%
Show abstract

Research replication is essential to build trust in evidence produced from real-world data. However, methods for conducting and reporting these studies are lacking, particularly related to data quality and fitness assessments. We replicated a single-center study from Children's Hospital of Atlanta in a multi-institutional learning network (PEDSnet) to evaluate the long-term effects of hydroxyurea in children with severe sickle cell disease (SS/S{beta}0 genotype). An AS-IS arm applied the original study's criteria with no major data quality adjustments, while a Data Fitness Enhanced (DFE) arm used systematic data fitness assessment to inform adjustments to cohort inclusion criteria and variable definitions; both arms then replicated the original study's primary analyses. Data quality checks in the DFE arm refined cohort criteria and improved hydroxyurea capture, drug era computation, and hematology specialist mapping. The DFE cohort produced average treatment effects with higher face validity and greater concordance with the original study (e.g., change in ED visits: -0.44 (CI -0.60, -0.26) versus -0.36 (CI -0.57, -0.16) in the original study) than the AS-IS cohort (-0.08 (CI -0.26, 0.09)), which yielded several implausible results. These findings show that superficially plausible cohort characteristics do not guarantee valid results without transparent, systematic data fitness assessment.

17
RedFuMOS: A novel approach for multi-omics and clinical data-driven patient stratification

De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.

2026-08-31 health informatics 10.64898/2026.08.26.26361415 medRxiv
Top 0.2%
4.3%
Show abstract

Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.

18
EpiKG2DAG: a Framework for Automated DAG Construction from Biomedical Text

DU, J.; Deng, G.

2026-08-11 health informatics 10.64898/2026.08.09.26360023 medRxiv
Top 0.2%
4.3%
Show abstract

While Directed Acyclic Graphs (DAGs) are essential for causal inference, their construction often relies on expert heuristics, which bypasses systematic evidence synthesis and creates a critical "evidence retrieval gap" in causal modeling. This study introduces EpiKG2DAG, a framework that supports evidence-anchored candidate DAG generation by transforming unstructured biomedical abstracts into structured epidemiological associations. We utilized DeepSeek-V3 to extract exposure-outcome association triplets from 189,266 abstracts and employed SapBERT for semantic normalization against UMLS concepts. The resulting Epidemiological Knowledge Graph (EpiKG) enables the automated identification of candidate confounders, mediators, and colliders based on graph-theoretic motifs and literature-derived evidence. A case study on COVID-19 and AKI demonstrates that the framework uncovers non-obvious confounders, such as air pollution, while ensuring evidence traceability. This work contributes to the field by mitigating the knowledge-acquisition bottleneck and providing a transparent, reproducible foundation for evidence-based causal modeling.

19
When can predictive uncertainty be trusted? A methodological evaluation in free-living wearable electrocardiogram signal-quality assessment

Tran, K. D.

2026-08-28 health informatics 10.64898/2026.08.25.26361304 medRxiv
Top 0.2%
4.2%
Show abstract

Uncertainty quantification is proposed as a safeguard for machine-learning systems in health-related signal analysis, but an uncertainty score is useful only if it behaves as a reliability signal. Free-living wearable electrocardiogram (ECG) signal-quality assessment provides a test bed because ambiguity, artifact, and acquisition shift can alter the relationship between confidence and correctness. This study evaluates predictive uncertainty under ambiguity, controlled corruption, and external distribution shift. 32,224 non-overlapping 10-s windows of synchronised single-lead ECG and three-axis accelerometry from 15 subjects in the Brno University of Technology ECG Quality Database were analysed. Two model families were compared: multinomial logistic regression and Classification and Regression Tree (CART), each progressing from a point estimate to a fixed-structure posterior and then a structure posterior. Expected conditional entropy and mutual information were evaluated as designated aleatoric and epistemic uncertainty measures, with max-softmax uncertainty as a confidence baseline. Validation covered error ranking, selective prediction, behavioural probes, posterior structural diversity, recorded-noise stress testing, and zero-shot external transfer. The logistic structure posterior retained an expected 8.5 of nine features and concentrated on near-complete masks, yielding little additional predictive diversity. Bayesian CART produced 221 distinct complete topologies among 238 retained draws and stronger score-dependent selective-risk behaviour. Conditional entropy increased with local class overlap, whereas mutual information increased when training information was reduced, although both showed cross-sensitivity. Under recorded noise, predicted quality severity changed more consistently than uncertainty, while external transfer preserved ordinal severity more reliably than uncertainty ordering. These findings show that posterior richness alone does not establish reliable uncertainty. Model-derived uncertainty should therefore be validated against prespecified ambiguity, information, and shift probes before supporting abstention, reacquisition, or downstream decisions.

20
Machine learning for elective caesarean section in Bangladesh: validation design, not model choice, determines the performance a deployed model would have

Rony, A. R.; Nahin, K. S. A.; Islam, T.; Asha, A. S.; Hossen, A.

2026-08-13 health informatics 10.64898/2026.08.12.26360275 medRxiv
Top 0.3%
4.2%
Show abstract

Caesarean section in Bangladesh reached 51.8% of deliveries in 2025, and elective caesarean, meaning caesarean before labour began, reached 31.6%. Risk models built on national household surveys are increasingly proposed for pointing audit toward places where scheduled surgery is outrunning clinical need, but they are usually validated in ways that flatter them. Using the 2025 Bangladesh Multiple Indicator Cluster Survey, we developed four models on 9,538 women (logistic regression, elastic net, random forest, gradient boosting) and ran the same procedure under three validation designs: random five-fold cross-validation; five-fold cross-validation grouped by sampling cluster; and leave-one-division-out cross-validation. We also tested transfer between the 2019 and 2025 rounds and audited subgroup calibration. No model improved on logistic regression by a margin worth acting on: the area under the receiver operating characteristic curve ranged from 0.724 to 0.736 under cluster-grouped validation, a spread of 0.012. Validation design mattered far more than the algorithm. Grouping folds by sampling cluster changed discrimination by at most 0.0004, this survey contributing a median of 3 eligible women per enumeration area. Withholding a whole division cost 0.044 to 0.060, more than 100 times as much, and still cost 0.033 to 0.056 after the strongest predictor, an outcome-derived district rate, was removed from every model. A model fitted to 2019 data lost 0.083 when applied to 2025, and the two rounds agreed only moderately on which predictors mattered (Spearman rank correlation 0.61). Calibration held in every wealth quintile, both residence categories and seven of eight divisions; Sylhet was the exception. Elective caesarean is predictable from routine survey items, but that predictability is local. Cross-validation, including cluster-aware cross-validation, does not measure what a model would do in a district it has never seen; a geographic holdout is the cheapest design that does.